Papers by Mohd Mujtaba Akhtar
HCFD: A Benchmark for Audio Deepfake Detection in Healthcare (2026.findings-acl)
Copied to clipboard
| Challenge: | a new task for detecting codec-fakes under pathological speech conditions is presented . we focus on codec based synthetic speech since neural codec decoding is a core building block in speech generation pipelines. |
| Approach: | They propose a new task for detecting codec-fakes under pathological speech conditions . they focus on codec based synthetic speech since neural codec decoding is a core building block in speech pipelines . |
| Outcome: | The proposed framework outperforms speech-based models on Healthcare CodecFake . it achieves the strongest performance on the task across clinical conditions and codecs . |
DIVINE : Coordinating Multimodal Disentangled Representations for Oro-Facial Neurological Disorder Assessment (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing frameworks for diagnosing oro-facial neurological disorders are based on shared and modality-specific representations, but they are not fully disentangled. |
| Approach: | They propose a fully disentangled multimodal framework that captures vocal and facial cues. |
| Outcome: | The proposed framework achieves 98.26% accuracy and 97.51% F1-score under modality-constrained scenarios. |
Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages (2026.findings-acl)
Copied to clipboard
| Challenge: | Speech deepfakes are highly realistic and can generate a few seconds of recorded speech. |
| Approach: | They propose an ALM that integrates semantic and prosodic representations from Whisper and TRILLsson to generate a speech deepfake dataset. |
| Outcome: | The proposed framework outperforms existing ALMs on the ICF benchmark in Indic languages. |
Prosody as Supervision: Bridging the Non-Verbal–Verbal for Multilingual Speech Emotion Recognition (2026.acl-long)
Copied to clipboard
| Challenge: | Existing paradigms for low-resource multilingual speech emotion recognition rely on labeled verbal speech and lack cross-lingual transfer. |
| Approach: | They propose a paralinguistic supervision paradigm for low-resource multilingual speech emotion recognition that leverages non-verbal vocalizations to exploit prosody-centric emotion cues. |
| Outcome: | The proposed framework outperforms Euclidean counter parts and strong SSL baselines in the language-based evaluation of low-resource multilingual speech emotion recognition (LRM-SER) |
Bridging Attribution and Open-Set Detection using Graph-Augmented Instance Learning in Synthetic Speech (2026.eacl-long)
Copied to clipboard
| Challenge: | Synthetic speech detection is a critical part of safeguarding digital communication, enabling systems to identify and mitigate the risks posed by highly realistic, machine-generated voices. |
| Approach: | They propose a framework that combines SFMs with graph-based modeling and open-set generalization to capture meaningful relationships between utterances and recognize speech that doesn’t belong to any known generator. |
| Outcome: | The proposed framework improves performance across both tasks, with Mamba-based embeddings delivering particularly strong results. |